Back

Mobile DNA

Springer Science and Business Media LLC

All preprints, ranked by how well they match Mobile DNA's content profile, based on 31 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Exploring Alu-Driven DNA Transductions in the Primate Genomes

Halabian, R.; Storer, J. M.; Hoyt, S. J.; Hartley, G. A.; Brosius, J.; O'Neill, R. J.; Makalowski, W.

2024-04-30 genetics 10.1101/2024.04.29.591526 medRxiv
Top 0.1%
77.8%
Show abstract

Long terminal repeats (LTRs) and non-LTRs retrotransposons, aka retroelements, collectively occupy a substantial part of the human genome. Certain non-LTR retroelements, such as L1 and SVA, have the potential for DNA transduction, which involves the concurrent mobilization of flanking non-transposon DNA during retrotransposition. These events can be detected by computational approaches. Despite being the most abundant short interspersed sequences (SINEs) that are still active within the genomes of humans and other primates, the transduction rate caused by Alu sequences remains unexplored. Therefore, we conducted an analysis to address this research gap and utilized an in-house program to probe for the presence of Alu-related transductions in the human genome. We analyzed 118,489 full-length AluY subfamilies annotated within the first complete human reference genome, T2T-CHM13. For comparative insights, we extended our exploration to two non-human primate genomes, the chimpanzee and the rhesus monkey. After manual curation, our findings did not confirm any Alu-mediated transductions, whose source genes are, unlike L1 or SVA, transcribed by RNA polymerase III, implying that they are infrequent or possibly absent not only in the human but also in chimpanzee and rhesus monkey genomes. Although we identified loci in which the 3 Target Site Duplication (TSD) was located distantly from the retrotransposed AluYs, a transduction hallmark, our study could not find further support for such events. The observation of these instances can be explained by the incorporation of other nucleotides into the poly(A) tails in conjunction with polymerase slippage.

2
Hide and seek: de novo identification in sugar beet reveals impact of non-autonomous LTR retrotransposons

Maiwald, S.; Maiwald, F.; Heitkam, T.

2026-03-03 genomics 10.64898/2026.03.01.708851 medRxiv
Top 0.1%
54.9%
Show abstract

Plant genomes are filled with retrotransposons and their derivatives, subject to constant sequence turnover. As short, non-autonomous retrotransposons do not encode a protein product, they experience reduced selective constraints on their DNA sequence, leading to diversification into multiple families, usually limited to only a few species. This absence of any coding capacity and their tendency to form subfamilies are the reasons for the incomplete description of non-autonomous LTR retrotransposons in most to all genomic repeat annotations. Here, we focus on non-autonomous LTR retrotransposon identification. Are all of these sequences derivatives of easier-to-identify full-length elements? Or is there more variability, which is currently overlooked? For this, we capitalize on our comprehensive understanding of the TE landscape in sugar beet to assess the extent of the blind spot on non-autonomous LTR retrotransposons Here, we present a workflow to identify non-autonomous LTR retrotransposons without prior sequence information, retrieving more than 100 families within the sugar beet genome. We only include TEs without the ability for complete self mobilization. Spanning up to 15,000 bp, these non-autonomous families are often longer than expected and characterized by reshuffling and modular evolution. Most strikingly, only a few of these families are directly derived from autonomous partners, showing that there is a large, undiscovered TE variety in the non-autonomous TE fraction. We highlight that a large fraction of non-autonomous TEs wont be retrieved with the current TE identification workflows, even if the output is well-curated and condensed into TE libraries and suggest procedures to remedy this gap. This study is the first insight into the non-autonomous LTR retrotransposon landscape within a single genome and serves as an example to estimate the error in non-autonomous TE detection.

3
Involvement of non-LTR retrotransposons in the cancer incidence and lifespan of mammals

Ricci, M.; Peona, V.; Taccioli, C.

2021-09-28 evolutionary biology 10.1101/2021.09.27.461867 medRxiv
Top 0.1%
54.9%
Show abstract

The presence in nature of closely related species showing drastic differences in lifespan and cancer incidence has recently increased the interest of the scientific community on these topics. In particular, the adaptations and genomic characteristics underlying the evolution of cancer-resistant and long-lived species have recently focused on the presence of alterations in the number of non-coding RNAs, on epigenetic regulation and, finally, on the activity of transposable elements (TEs). In this study, we compared the content and dynamics of TE activity in the genomes of four rodent and six bat species exhibiting different lifespans and cancer susceptibility. Mouse, rat and guinea pig (short-lived and cancer-prone organisms) were compared with the naked mole rat (Heterocephalus glaber) which is the rodent with the longest lifespan. The long-lived and cancer-resistant bats of the genera Myotis, Rhinolophus, Pteropus and Rousettus were instead compared with the Molossus, which is instead a short-lived and cancer-resistant organism. Analyzing the patterns of recent accumulations of TEs in the genome in these species, we found a strong suppression or negative selection to accumulation, of non-LTR retrotransposons in long-lived and cancer-resistant organisms. On the other hand, all short-lived and cancer-prone species have shown recent accumulation of this class of TEs. Among bats, the Molossus molossus turned out to be a very particular species and, at the same time, an important model because, despite being susceptible to rapid ageing, it is resistant to cancer. In particular, we found that its genome has the highest density of SINE (non-LTR retrotransposons), but, on the other hand, a total lack of active LINE retrotransposons. Our hypothesis is that the lack of LINEs presumably makes the Molossus cancer resistant due to lack of retrotransposition but, at the same time, the high presence of SINE, may be related to their short life span due to "sterile inflammation" and high mutation load. We suggest that research on ageing and cancer evolution should put particular attention to the involvement of non-LTR retrotransposons in these phenomena.

4
Jumping between Turtles, Fishes, and a Frog: The Unexpected Horizontal Transfer of a DNA Transposon

Hassan, N. T.; Adelson, D. L.; Galbraith, J. D.

2023-08-21 evolutionary biology 10.1101/2023.08.19.550906 medRxiv
Top 0.1%
42.1%
Show abstract

Horizontal transfer of transposable elements (HTT) has been reported across many species and the impact of such events on genome structure and function has been well described. However, few studies have focused on reptilian genomes, especially HTT events in Testudines (turtles). Here, we investigated the repetitive content of Malaclemys terrapin terrapin (Diamondback turtle) and found a high similarity hAT-6 DNA transposon shared between other turtle species, ray-finned fishes, and a frog. hAT-6 was notably absent in taxa closely related to turtles, such as crocodiles and birds. Successful invasion of DNA transposons into new genomes requires the conservation of specific residues in the encoded transposase, and through structural analysis, these residues were identified indicating retention of functional transposition activity. We document a rare and recent HTT event of a DNA transposon between turtles which are known to have a low genomic evolutionary rate and ancient repeats.

5
DNA satellite and chromatin organization at house mouse centromeres and pericentromeres

Packiaraj, J.; Thakur, J.

2023-07-19 genomics 10.1101/2023.07.18.549612 medRxiv
Top 0.1%
40.4%
Show abstract

Centromeres are essential for faithful chromosome segregation during mitosis and meiosis. However, the organization of satellite DNA and chromatin at mouse centromeres and pericentromeres is poorly understood due to the challenges of sequencing and assembling repetitive genomic regions. Using recently available PacBio long-read sequencing data from the C57BL/6 strain and chromatin profiling, we found that contrary to the previous reports of their highly homogeneous nature, centromeric and pericentromeric satellites display varied sequences and organization. We find that both centromeric minor satellites and pericentromeric major satellites exhibited sequence variations within and between arrays. While most arrays are continuous, a significant fraction is interspersed with non-satellite sequences, including transposable elements. Additionally, we investigated CENP-A and H3K9me3 chromatin organization at centromeres and pericentromeres using Chromatin immunoprecipitation sequencing (ChIP-seq). We found that the occupancy of CENP-A and H3K9me3 chromatin at centromeric and pericentric regions, respectively, is associated with increased sequence abundance and homogeneity at these regions. Furthermore, the transposable elements at centromeric regions are not part of functional centromeres as they lack CENP-A enrichment. Finally, we found that while H3K9me3 nucleosomes display a well-phased organization on major satellite arrays, CENP-A nucleosomes on minor satellite arrays lack phased organization. Interestingly, the homogeneous class of major satellites phase CENP-A and H3K27me3 nucleosomes as well, indicating that the nucleosome phasing is an inherent property of homogeneous major satellites. Overall, our findings reveal that house mouse centromeres and pericentromeres, which were previously thought to be highly homogenous, display significant diversity in satellite sequence, organization, and chromatin structure.

6
SINE Retrotransposons Import Polyadenylation Signals to 3'UTRs in Dog (Canis familiaris)

Choi, J. D.; Del Pinto, L. A.; Sutter, N. B.

2020-12-01 genomics 10.1101/2020.11.30.405357 medRxiv
Top 0.1%
38.7%
Show abstract

BackgroundMessenger RNA 3 untranslated regions (3UTRs) control many aspects of gene expression and determine where the transcript will terminate. The polyadenylation signal (PAS) AAUAAA is a key regulator of transcript termination and this hexamer, or a similar sequence, is very frequently found within 30 bp of 3UTR ends. Short interspersed element (SINE) retrotransposons are found throughout genomes in high copy number. When inserted into genes they can disrupt expression, alter splicing, or cause nuclear retention of mRNAs. The genomes of the domestic dog and other carnivores carry hundreds of thousands Can-SINEs, a tRNA-related SINE with transcription termination potential. Because of this we asked whether Can-SINEs may help terminate transcript in some dog genes. ResultsDog 3UTRs have several peaks of AATAAA PAS frequency within 40 bp of the 3UTR end, including four bp-interval peaks at 28, 32, and 36 bp from the end. The periodicity is partly explained by TAAA(n) repeats within Can-SINE AT-rich tails. While density of antisense-oriented Can-SINEs in 3UTRs is fairly constant with distances from 3end, sense-oriented Can-SINEs are common at the 3end but nearly absent farther upstream. There are nine Can-SINE sub-types in the dog genome and the consensus sequence sense strands (head to tail) all carry at least three PASs while antisense strands usually have none. We annotated all repeat-masked Can-SINE copies in the Boxer reference genome and found that the young SINEC_Cf type has a mode of 15 bp for target site duplications (TSDs). We find that all Can-SINE types favor integration at TSDs beginning with A(4). The count of AATAAA PASs differs significantly between sense and antisense-oriented retrotransposons in transcripts. Can-SINEs near 3UTR ends are very likely to carry AATAAA on the mRNA sense strand while those farther upstream are not. We also identified loci where Can-SINE insertion has truncated or altered a dog 3UTR compared to the human ortholog. ConclusionDog Can-SINE activity has imported AATAAA PASs into gene transcripts and led to alteration of 3UTRs. AATAAA sequences are selectively removed from Can-SINEs in introns and upstream 3UTR regions but are retained at the far downstream end of 3UTRs, which we infer reflects their role as termination sequences for these transcripts.

7
FREDDIE: A comprehensive tool for detecting exonization of retrotransposable elements in short and long RNA sequencing data

Mercuri, R. L. V.; Miller, T. L. A.; dos Santos, F. F.; de Lima, M. F.; Rangel-Pozzo, A.; Galante, P. A. F.

2024-04-26 bioinformatics 10.1101/2024.04.22.590610 medRxiv
Top 0.1%
30.4%
Show abstract

BackgroundTransposable elements (TEs) constitute a significant portion of mammalian genomes, accounting for about 50% of the total DNA. Intragenic TEs are of particular interest as they are co-transcribed with their host genes in pre-mRNA, potentially leading to the formation of novel chimeric transcripts and the exonization of TEs. The abundance of RNA sequencing data currently available offers a unique opportunity to explore transcriptomic variations. However, a significant limitation is the capability of existing computational tools. Here, we introduce FREDDIE, an innovative algorithm designed to detect the exonization of retrotransposable elements using RNA-seq data. FREDDIE can process short and long RNA sequencing data, assemble and quantify transcripts, evaluate coding potential, and identify protein domains in chimeric transcripts involving exonized TEs and retrocopies. ResultsTo demonstrate the efficacy of FREDDIE, we analyzed and validated TE exonization in two human cancer cell lines, K562 and U251. We have identified 322 chimeric transcripts, of which 126 were from K562, and 196 were from U251. Among these chimeric transcripts, there were 35 that showed similar exonization patterns and host genes. These transcripts involve protein-coding genes of the host and exonization of LINE-1 (L1), Alu elements, and retrocopies of coding genes. We have selected some candidates and validated them experimentally through RT-PCR. The validation rate for these candidates was 70%, later confirmed by long-read sequencing. Additionally, we applied FREDDIE to analyze TE exonization across 157 glioblastoma samples, identifying 1,010 chimeric transcripts. The majority of these transcripts involved the exonization of Alu elements (69.8%), followed by L1 (20.6%) and retrocopies (9.6%). Notably, we discovered a highly expressed L1 exonization within the ROS gene, resulting in a truncated open reading frame (ORF) with the deletion of two protein domains. ConclusionsFREDDIE is an efficient and user-friendly tool for identifying chimeric transcripts that involve exonization of intragenic TEs. Overall, FREDDIE enables comprehensive investigations into the contributions of TEs to transcriptome evolution, variation, and disease-associated abnormalities, and it operates effectively on standard computing systems. FREDDIE is publicly available: https://github.com/galantelab/freddie

8
Recurrent LINE 1 exonization drives transcriptome remodelling in NSCLC

Parida, A. S.; Kumar, A.; Tiwari, B.

2026-04-24 genomics 10.64898/2026.04.22.720055 medRxiv
Top 0.1%
29.7%
Show abstract

The only autonomously active transposable elements in the human genome are Long interspersed nuclear element-1 (LINE-1) elements. These elements are known to play an important role in changing the transcriptome. LINE-1 sequences affect gene regulation during post-transcription processing, along with their established role in retrotransposition. Exonization is one mechanism where the LINE-1 integrated genome undergoes alternative splicing to produce new isoforms of transcripts. Our work mainly highlights the effect of LINE-1 associated exonization, focusing on the formation of isoforms of transcripts. Using Non-small cell lung cancer (NSCLC) as a model, we conducted a detailed transcriptome study that combines splice junction profiling with gene expression data. Our results show that LINE-1 sequences are often included as exons in host transcripts, leading to the formation of new exons and their various isoforms. The events are validated by solid splice junction evidence that proves the reliability and reproducibility. In particular, it was observed that repetitive analyses revealed certain LINE-1 exonization events that were consistent. The finding indicates that LINE-1 act as recurrent sources of splice ready sequences. Though exonizations do not necessarily affect the total expression levels of genes, our study reveals that they certainly contribute to transcript diversity. The diversity of isoforms generated potentially contributes to the effects of gene function. This study is limited to NSCLC, but it is likely that the exonizations events play a crucial role in the altering RNA diversity in cancers. Therefore the study elucidates new insights into how transposable elements modify gene structure and function during cancer development.

9
Localized assembly for long reads enables genome-wide analysis of repetitive regions at single-base resolution in human genomes

Ikemoto, K.; Fujimoto, H.; Fujimoto, A.

2022-12-03 genomics 10.1101/2022.12.02.518938 medRxiv
Top 0.1%
26.7%
Show abstract

BackgroundLong-read sequencing technologies have the potential to overcome the limitations of short reads and provide a comprehensive picture of the human genome. However, it remains hard to characterize repetitive sequences by reconstructing genomic structures at high resolution solely from long reads. Here, we developed a localized assembly method (LoMA) that constructs highly accurate consensus sequences (CSs) from long reads. MethodsWe first developed LoMA, by combining minimap2, MAFFT, and our algorithm, which classifies diploid haplotypes based on structural variants and constructs CSs. Using this tool, we analyzed two human samples (NA18943 and NA19240) sequenced with the Oxford Nanopore sequencer. We defined target regions in each genome based on mapping patterns and then constructed a high-quality catalog of the human insertion solely from the long-read data. ResultsThe assessment of LoMA showed high accuracy of CSs (error rate < 0.3%) compared with raw data (error rate > 8%) and superiority to the previous study. The genome-wide analysis of NA18943 and NA19240 identified 5,516 and 6,542 insertions ({zeta} 100 bp) respectively. Most insertions ([~]80%) were derived from the tandem repeat and transposable elements. We also detected processed pseudogenes, insertions in transposable elements, and long insertions (> 10 kbp). Further, our analysis suggested that short tandem duplications were association with gene expression and transposons. ConclusionsOur analysis showed that LoMA constructs high-quality sequences from long reads with substantial errors. This study revealed the true structures of insertions with high accuracy and inferred mechanisms for the insertions. Our approach contributes to the future human genome studies. LoMA is available at our GitHub page: https://github.com/kolikem/loma.

10
Rapid evolutionary diversification of the flamenco locus in the D. simulans clade

Signor, S.; Vedanayagam, J.; Kim, B. Y.; Wierzbicki, F.; Kofler, R.; Lai, E. C.

2022-11-17 evolutionary biology 10.1101/2022.09.29.510127 medRxiv
Top 0.1%
26.5%
Show abstract

Effective suppression of transposable elements (TEs) is paramount to maintain genomic integrity and organismal fitness. In D. melanogaster, flamenco is a master suppressor of TEs, preventing their movement from somatic ovarian support cells to the germline. It is transcribed by Pol II as a long (100s of kb), single-stranded, primary transcript, that is metabolized into Piwi-interacting RNAs (piRNAs) that target active TEs via antisense complementarity. flamenco is thought to operate as a trap, owing to its high content of recent horizontally transferred TEs that are enriched in antisense orientation. Using newly-generated long read genome data, which is critical for accurate assembly of repetitive sequences, we find that flamenco has undergone radical transformations in sequence content and even copy number across simulans clade Drosophilid species. D. simulans flamenco has duplicated and diverged, and neither copy exhibits synteny with D. melanogaster beyond the core promoter. Moreover, flamenco organization is highly variable across D. simulans individuals. Next, we find that D. simulans and D. mauritiana flamenco display signatures of a dual-stranded cluster, with ping-pong signals in the testis and/or embryo. This is accompanied by increased copy numbers of germline TEs, consistent with these regions operating as functional dual stranded clusters. Overall, the physical and functional diversity of flamenco orthologs is testament to the extremely dynamic consequences of TE arms races on genome organization, not only amongst highly related species, but even amongst individuals.

11
Regulatory Features and Functional Specialization of Human Endogenous Retroviral LTRs: A Genome-Wide Annotation and Analysis via HERVarium

Montserrat-Ayuso, T.; Pujol, A.; Esteve-Codina, A.

2026-02-18 genomics 10.64898/2026.02.17.706328 medRxiv
Top 0.1%
26.5%
Show abstract

Human endogenous retroviruses (HERVs), remnants of ancient retroviral infections, account for over 8% of the human genome and remain an underexplored source of cis-regulatory elements and protein-coding remnants. Here we present HERVarium, an integrated and interactive database that provides access to systematic annotations of the protein-domain architecture of internal HERV regions and the regulatory landscape of their LTRs across the human genome. Building on our previous catalog of conserved retroviral domains, we classified over 400,000 LTRs by structure (solo, 5', or 3' LTR), proximity to transcription start sites (TSS), potential transcription-factor binding-motif (TFBM) burden, and computationally reconstructed their canonical U3-R-U5 substructure, revealing that two-thirds retain recognizable segmentation, particularly those adjacent to conserved internal domains. LTRs flanking internal regions with conserved Gag, Pol, or Env domains tend to be longer and richer in motifs, suggesting coordinated maintenance of coding and regulatory potential. Conversely, solo LTRs positioned at gene TSSs exhibit promoter-like architectures enriched in motifs for developmental and proliferative regulators, while selectively lacking motifs associated with neuronal differentiation. Some 5' LTRs at lncRNA TSSs also flank conserved retroviral domains, indicating that certain lncRNAs may derive from transcriptionally active, structurally intact HERV loci. HERVarium provides a comprehensive, interactive resource to explore and download these annotations at single-locus resolution.

12
Identification and characterization of retro-DNAs, a new type of retrotransposons originated from DNA transposons, in primate genomes

Tang, W.; Liang, P.

2020-03-20 evolutionary biology 10.1101/2020.03.19.999144 medRxiv
Top 0.1%
22.9%
Show abstract

Mobile elements (MEs) can be divided into two major classes based on their transposition mechanisms as retrotransposons and DNA transposons. DNA transposons move in the genomes directly in the form of DNA in a cut-and-paste style, while retrotransposons utilize an RNA-intermediate to transpose in a "copy-and-paste" fashion. In addition to the target site duplications (TSDs), a hallmark of transposition shared by both classes, the DNA transposons also carry terminal inverted repeats (TIRs). DNA transposons constitute ~3% of primate genomes and they are thought to be inactive in the recent primate genomes since ~37My ago despite their success during early primate evolution. Retrotransposons can be further divided into Long Terminal Repeat retrotransposons (LTRs), which are characterized by the presence of LTRs at the two ends, and non-LTRs, which lack LTRs. In the primate genomes, LTRs constitute ~9% of genomes and have a low level of ongoing activity, while non-LTR retrotransposons represent the major types of MEs, contributing to ~37% of the genomes with some members being very young and currently active in retrotransposition. The four known types of non-LTR retrotransposons include LINEs, SINEs, SVAs, and processed pseudogenes, all characterized by the presence of a polyA tail and TSDs, which mostly range from 8 to 15 bp in length. All non-LTR retrotransposons are known to utilize the L1-based target-primed reverse transcription (TPRT) machineries for retrotransposition. In this study, we report a new type of non-LTR retrotransposon, which we named as retro-DNAs, to represent DNA transposons by sequence but non-LTR retrotransposons by the transposition mechanism in the recent primate genomes. By using a bioinformatics comparative genomics approach, we identified a total of 1,750 retro-DNAs, which represent 748 unique insertion events in the human genome and nine non-human primate genomes from the ape and monkey groups. These retro-DNAs, mostly as fragments of full-length DNA transposons, carry no TIRs but longer TSDs with ~23.5% also carrying a polyA tail and with their insertion site motifs and TSD length pattern characteristic of non-LTR retrotransposons. These features suggest that these retro-DNAs are DNA transposon sequences likely mobilized by the TPRT mechanism. Further, at least 40% of these retro-DNAs locate to genic regions, presenting significant potentials for impacting gene function. More interestingly, some retro-DNAs, as well as their parent sites, show certain levels of current transcriptional expression, suggesting that they have the potential to create more retro-DNAs in the current primate genomes. The identification of retro-DNAs, despite small in number, reveals a new mechanism in propagating the DNA transposons sequences in the primate genomes with the absence of canonical DNA transposon activity. It also suggests that the L1 TPRT machinery may have the ability to retrotranspose a wider variety of DNA sequences than what we currently know.

13
Transposable element competition in shaping the deer mouse genome

Gozashti, L.

2022-10-20 genomics 10.1101/2022.10.18.512801 medRxiv
Top 0.1%
22.3%
Show abstract

The genomic landscape of transposable elements (TEs) varies dramatically across species, with some TEs demonstrating greater success in colonizing particular lineages than others. In mammals, LINE retrotransposons typically occupy more of the genome than any other TE and most LINE content is represented by a single family: L1. Here, we report an unusual genomic landscape of TEs in the deer mouse, Peromyscus maniculatus, a model for studying the genomic basis of adaptation. In contrast to other previously examined mammalian species, LTR elements occupy more of the deer mouse genome than LINEs (11% and 10% respectively). This pattern reflects a combination of relatively low LINE activity in addition to a massive invasion of lineage-specific endogenous retroviruses (ERVs). Deer mouse ERVs exhibit diverse origins spanning the retroviral phylogeny suggesting that these rodents have been host to a wide range of exogenous retroviruses. Notably, we were able to trace the origin of one ERV lineage, which arose within the last [~]11-18 million years, to a close relative of feline leukemia virus, revealing inter-ordinal horizontal transmission of these zoonotic viruses. Several lineage-specific ERV subfamilies have attained very high copy numbers, with the top five most abundant accounting for [~]2% of the genome. Concomitant to the expansive diversification of ERVs, we also observe a massive expansion of Kruppel-associated box domain-containing zinc finger genes (KZNFs), which likely control ERV activity and whose expansion may have been partially facilitated by ectopic recombination between ERVs. We also find evidence that ERVs directly impacted the evolutionary trajectory of LINEs by outcompeting them for genomic sites and frequently disrupting autonomous LINE copies. Together, our results illuminate the genomic ecology that shaped the deer mouse genomes TE landscape, opening up a range of opportunities to investigate the evolutionary processes that give rise to variation in mammalian genome structure. SummaryTransposable elements (TEs) are a highly diverse collection of genetic elements capable of mobilizing in genomes and function as important drivers of genome evolution. The landscape of TEs in a genome have been compared to a genomic ecosystem, with interactions between TEs and each other as well as TEs and their host, dictating the evolutionary success of TE lineages. While TE diversity and copy numbers can vary dramatically across taxa, the evolutionary reasons for this variation remain poorly understood. In mammals, long interspersed nuclear elements (LINEs) typically dominate, occupying more of the genome than any other TE. Here, we report a unique case in the deer mouse (Peromyscus maniculatus) in which long terminal repeat (LTR) retrotransposons occupy more of the genome than LINEs. We investigate the evolutionary origins and implications of the deer mouses distinct genomic landscape, revealing ecological processes that helped shape its evolution. Together, our results provide much-needed insight into the evolutionary processes that give rise to variation in mammalian genome structure.

14
Functional validation of transposable element derived cis-regulatory elements in Atlantic salmon

Sahlstrom, H. M.; Datsomor, A. K.; Monsen, O.; Hvidsten, T. R.; Sandve, S. R.

2022-11-03 evolutionary biology 10.1101/2022.11.02.514921 medRxiv
Top 0.1%
21.8%
Show abstract

BackgroundTransposable elements (TEs) are hypothesized to play important roles in shaping genome evolution following whole genome duplications (WGD), including rewiring of gene regulation. In a recent analysis, duplicate gene copies that had evolved higher expression in liver following the salmonid WGD ~100 million years ago were associated with higher numbers of predicted TE-derived cis-regulatory elements (TE-CREs). Yet, the ability of these TE-CREs to recruit transcription factors (TFs) in vivo and impact gene expression remains unknown. ResultsHere, we evaluated the gene regulatory functions of 11 TEs using luciferase promoter reporter assays in Atlantic salmon (Salmo salar) primary liver cells. Canonical Tc1-Mariner elements from intronic regions showed no or small repressive effects on transcription. However, other TE-derived cis-regulatory elements upstream of transcriptional start sites increased expression significantly. ConclusionOur results question the hypothesis that TEs in the Tc1-Mariner superfamily, which were extremely active following WGD in salmonids, had a major impact on regulatory rewiring of gene duplicates, but highlights the potential of other TEs in post-WGD rewiring of gene regulation in the Atlantic salmon genome.

15
Transposable element dynamics are consistent across the Drosophila phylogeny, despite drastically differing content

Hill, T.

2019-07-26 genomics 10.1101/651059 medRxiv
Top 0.1%
21.3%
Show abstract

BackgroundThe evolutionary dynamics of transposable elements (TEs) vary across the tree of life and even between closely related species with similar ecologies. In Drosophila, most of the focus on TE dynamics has been completed in Drosophila melanogaster and the overall pattern indicates that TEs show an excess of low frequency insertions, consistent with their frequent turn over and high fitness cost in the genome. Outside of D. melanogaster, insertions in the species Drosophila algonquin, suggests that this situation may not be universal, even within Drosophila. Here we test whether the pattern observed in D. melanogaster is similar across five Drosophila species that share a common ancestor more than fifty million years ago.\n\nResultsFor the most part, TE family and order insertion frequency patterns are broadly conserved between species, supporting the idea that TEs have invaded species recently, are mostly costly and dynamics are conserved in orthologous regions of the host genome\n\nConclusionsMost TEs retain similar activities and fitness costs across the Drosophila phylogeny, suggesting little evidence of drift in the dynamics of TEs across the phylogeny, and that most TEs have invaded species recently.

16
Genomic Environments and Their Influence on Transposable Element Communities

Saylor, B.; Kremer, S. C.; Gregory, T. R.; Cottenie, K.

2019-06-11 genomics 10.1101/667121 medRxiv
Top 0.1%
18.7%
Show abstract

BackgroundDespite decades of research the factors that cause differences in transposable element (TE) distribution and abundance within and between genomes are still unclear. Transposon Ecology is a new field of TE research that promises to aid our understanding of this often-large part of the genome by treating TEs as species within their genomic environment, allowing the use of methods from ecology on genomic TE data. Community ecology methods are particularly well suited for application to TEs as they commonly ask questions about how diversity and abundance of a community of species is determined by the local environment of that community.\n\nResultsUsing a redundancy analysis, we found that ~ 50% of the TEs within a diverse set of genomes are distributed in a predictable pattern along the chromosome, and the specific TE superfamilies that show these patterns are relate to the phylogeny of the host taxa. In a more focused analysis, we found that ~60% of the variation in the TE community within the human genome is explained by its location along the chromosome, and of that variation two thirds (~40% total) was explained by the 3D location of that TE community within the genome (i.e. what other strands of DNA physically close in the nucleus). Of the variation explained by 3D location half (20% total) was explained by the type of regulatory environment (sub compartment) that TE community was located in. Using an analysis to find indicator species, we found that some TEs could be used as predictors of the environment (sub compartment type) in which they were found; however, this relationship did not hold across different chromosomes.\n\nConclusionsThese analyses demonstrated that TEs are non-randomly distributed across many diverse genomes and were able to identify the specific TE superfamilies that were non-randomly distributed in each genome. Furthermore, going beyond the one-dimensional representation of the genome as a linear sequence was important to understand TE patterns within the genome. Additionally, we extended the utility of traditional community ecology methods in analyzing patterns of TE diversity.

17
Introns are derived from transposons

Rogers, S. O.; Bendich, A. J.

2023-02-22 evolutionary biology 10.1101/2023.02.21.529479 medRxiv
Top 0.1%
18.4%
Show abstract

Introns and transposons exhibit many similar features, but the connections between them have yet to be firmly established. Group I introns have commonalities with DNA transposons, while group II introns share many features with retrotransposons. Here, we report the results of an analysis of 214 introns (including group I, group II, group III, twintrons, spliceosomal, and archaeal introns) from members of seven major taxa (within Eukarya, Bacteria, and Archaea) that all have direct repeats at or near both exon/intron borders, indicating that they were inserted via transposition events. Border sequence analysis indicates that after splicing, most mature transcripts would be functionally compromised because they do not restore the DNA sequence information before intron insertion. Transposons and introns thus appear to be members of a diverse assemblage of parasitic mobile genetic elements that secondarily may benefit their host cell and have expanded greatly in eukaryotes from their presumed prokaryotic ancestors. Author SummaryIntrons are found in all domains of life. While they are limited in prokaryotes, they have greatly expanded in number and diversity in eukaryotes. We found direct repeat sequences at or near both exon/intron borders for all 214 introns analyzed among eukaryotes, bacteria, and archaea. We infer that all introns were inserted into genes via transposon-like mechanisms and are members of a large family of mobile genetic elements.

18
NeuralTE: an accurate approach for Transposable Element superfamily classification with multi-feature fusion

Hu, K.; Xu, M.; Gao, X.; Wang, J.

2024-04-28 bioinformatics 10.1101/2024.01.21.576519 medRxiv
Top 0.1%
18.4%
Show abstract

MotivationClassifying Transposable Elements (TEs) at the superfamily level offers deeper insights into species variation and evolution. Recent advancements in third-generation sequencing technologies have made a large number of genomes from non-model species becoming available. However, existing TE classification methods suffer from several limitations, including the necessity to train multiple hierarchical classification models, the incapacity to perform classification at the superfamily level, and deficiencies in both accuracy and robustness. Therefore, there is an urgent need for an accurate TE classification method to improve genome annotation. ResultsIn this study, we develop NeuralTE, a deep learning method designed to classify transposons at the superfamily level. To achieve accurate TE classification, we identify various structural features of transposons, and use different combinations of k-mers for terminal repeats and internal sequences to uncover distinct patterns. Evaluation on all transposons from Repbase shows that NeuralTE outperforms existing deep learning, machine learning, and homology-based methods in classifying TEs. Testing on the transposons from novel species highlights the superior performance of NeuralTE compared to existing methods. We also conduct TE annotation experiments on rice using different classification tools, and the results show that NeuralTE achieves annotations nearly identical to the gold standard, highlighting its robustness and accuracy in classifying transposons. AvailabilityNeuralTE is publicly available at https://github.com/CSU-KangHu/NeuralTE.

19
Active and unusually expanded PIF/Harbinger transposable elements in the Caenorhabditis inopinata genome

Sato, K.; Jin, X.; Oomura, S.; Kawahara, K.; Sun, S.; Yoshida, A.; Haruta, N.; Sugimoto, A.; Kikuchi, T.

2026-05-29 genetics 10.64898/2026.05.26.728016 medRxiv
Top 0.1%
18.4%
Show abstract

BackgroundTransposable elements (TEs) serve as powerful drivers of genome innovation but also threaten genome integrity. The PIF/Harbinger superfamily is distinctive among DNA transposons because mobilisation typically requires proteins, a DDE transposase and a MADF DNA-binding protein. Caenorhabditis inopinata, the closest known relative of C. elegans, has a TE-rich genome and lacks multiple components of the ERGO-1-class endogenous small-RNA pathway, making it a useful system for examining TE dynamics in a distinct host context. We identified a spontaneous dumpy mutant of C. inopinata caused by insertion of a PIF/Harbinger-family element into the coding region of Cin-dpy-11. The inserted element, designated Harbinger-1M_cIno, belongs to the Turmoil2 lineage originally defined in C. elegans and retains a MADF domain but lacks a recognisable DDE transposase ORF. Genome-wide curation recovered 258 related copies, revealing a strongly asymmetric family structure. Short noncoding derivatives were predominant, MADF-bearing derivatives were expanded and only one DDE-bearing locus retained an apparently intact transposase gene, suggesting that DDE and MADF functions are partitioned across distinct elements and may be supplied in trans during mobilisation. We also identified a second PIF/Harbinger-derived family, Harbinger-2M_cIno, associated with the Turmoil1 lineage. This family comprises 1,376 copies and therefore records substantial past amplification, but it lacks a detectable DDE source, shows greater sequence divergence and more degraded terminal structures than Harbinger-1M_cIno. Together, these data indicate that the two PIF/Harbinger lineages in C. inopinata differ not in whether amplification occurred, but in when it occurred and whether present-day mobilisation competence has been retained.

20
Recursive Repeat Extender (RRE): A recursive approach to automatically extend repeat element models

Falcon, F.; Tanaka, E. M.; Rodriguez-Terrones, D.

2026-04-17 bioinformatics 10.64898/2026.04.14.718546 medRxiv
Top 0.1%
18.3%
Show abstract

Repetitive elements, including transposable elements (TEs), are integral structural components of eukaryotic genomes; consequently, their identification and classification are crucial to their study. Several approaches have been developed to perform de novo genome-wide repeat identification through pairwise sequence comparisons; however, they often generate truncated repeat models due to their sampling strategies and the substantial fragmentation of many of the older repeat copies in the genome. To improve repeat models generated de novo, several algorithms have been developed that increase model length via the BEEA (BLAST-Extend-Extract-Align) approach, in which genomic instances of each repeat are identified with BLAST, their coordinates are extended, and a refined model is generated by aligning the extended sequences. Nevertheless, these extension algorithms exhibit two key limitations that hinder the reconstruction of highly degenerate and fragmented repeats: the use of BLAST as a search algorithm - which limits their sensitivity in detecting highly diverged sequences - and the use of a single search step, which precludes the reconstruction of extensively fragmented repeat models. In this work, we present a novel approach to extend repeat models, called RRE (Recursive Repeat Extender), which uses profile hidden Markov models (HMMs) to search for repeat elements with high sensitivity and employs a recursive extension strategy that iteratively searches and extends the repeat model, using the extended model from each round as input for the next and continuing until no additional sequence can be incorporated. We apply RRE to repeat libraries generated de novo from five model organisms, and our results show that RRE-generated repeat libraries contain fewer but longer repeat models and can identify a larger proportion of the genomes as repetitive than RepeatModeler2-generated repeat libraries. Notably, RRE can reconstruct highly degenerate repeats such as CR1_Mam, producing a model that achieves similar coverage to the reference Dfam model while extending it by an additional 131 bp that were not captured in the reference model. Overall, RRE enables the automatic improvement of de novo repeat libraries and the reconstruction of highly degenerate and fragmented repeats.